Tags: zero-shot detection*

0 bookmark(s) - Sort by: Date ↓ / Title /

  1. Ruoqi Guo et al. present RLCDAlignBench, a new benchmark designed to evaluate the ability of Jev—a model trained via reinforcement learning for calibrated decisions (RLCD)—to detect various types of alignment failures in language models zero-shot. The study examines ten specific failure modes, including sycophancy, jailbreaks, and hallucination, across 44 benchmarks and five target models. Results indicate that a single generic question applied to Jev achieves a median AUROC of 0.886, outperforming supervised baselines in many cases while being significantly more cost-effective than using LLM-judge scorers.
    - Evaluates ten failure modes: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking.
    - Uses a "relational" detection strategy by varying the question wording separately from input fields to handle failures defined against external references.
    - Jev costs 63x less than traditional LLM-judge scorers.

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: tagged with "zero-shot detection"

About - Propulsed by SemanticScuttle